Papers with Language State Tracker
OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing video dialog models struggle with questions requiring both spatial and temporal localization within videos, long-term temporal reasoning, and accurate object tracking across multiple dialog turns. |
| Approach: | They propose a multi-modal attention-based model for video dialog operating over a dialog state tracker. |
| Outcome: | The proposed model can learn multi-modal dialog state representations of the most relevant objects and rounds. |